Psychological Review
● American Psychological Association (APA)
Preprints posted in the last 30 days, ranked by how well they match Psychological Review's content profile, based on 19 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Ferrera, V. P.; Lippl, S.; Kay, K.; Munoz, F.; Jin, Y.; Jensen, G.; Terrace, H.
Show abstract
Transitive inference (TI) is the ability to reason about transitive relationships in an ordered set of items (e.g., if A>B and B>C, then A>C). TI is widely held to depend on a linear representation of the serial (rank) order of those items. By what computational mechanism is such an ordering constructed during learning, and how is it used to make choices that obey transitivity? Here we take a minimalist approach, applying least-squares estimation (LSE) to a serial learning task commonly used to test TI in humans and animals. In this formulation, LSE computes a linear classifier that maps task conditions onto behavioral outcomes. This algorithm makes no explicit assumptions about transitivity or serial order, yet it reproduces key empirical features of TI; namely, the ability to generalize beyond the training set, and a symbolic distance effect (SDE) in performance accuracy. Applying the classifier to individual items produces an internally ordered representation of rank from which both generalization and the SDE naturally emerge. The approach also yields a decision mechanism, in the form of a differencing operation, for selecting the correct item from any pair. These findings reframe TI as a linear classification problem, challenging conventional assumptions about the cognitive mechanisms required for transitive reasoning.
Kaltenmaier, A.; Press, C.
Show abstract
Past sensory experience shapes our perceptual decision-making in the now. Popular models frame perceptual decisions as either attracted towards or repelled away from recent sensory information, but it is unclear when and why these distinct effects emerge. We here ask whether effects turn from attractive to repulsive depending on the level of surprise elicited by the precision-weighted discrepancy between past and present sensory states. This model is based upon the idea that attraction is adaptive for optimizing efficiency and accuracy when discrepancies are small, because they likely reflect sensory noise rather than real change in the environment. In contrast, repulsion may reflect the upweighting of counterfactual evidence when discrepancies are large because they more likely signal the need for model updating. We test this model on a large amount of recently-collated trial-by-trial serial dependence data and consistently find support for it across the dataset, participant, and trial-by-trial level. Specifically, serial dependence effects are attractive at low discrepancies between past and current sensory states but turn repulsive when discrepancies are larger. Higher sensory precision is found to accelerate this flip by reducing the modal discrepancy threshold required to trigger repulsion effects. We discuss how these findings necessitate extending existing theories of serial dependence, and how they may resolve conflicts in the broader predictive processing, learning and perception literatures.
Li, M.; Jensen, K. T.; Zhang, Q.; Lu, Q.; Mattar, M. G.
Show abstract
Humans exhibit structured patterns of memory recall, including a tendency to recall more recent information and to recall events in the same order they were experienced. Classic computational models explain these patterns by positing that memories incorporate the ongoing ''temporal context'', formed by smoothly integrating the stimulus history. However, it is unclear whether a single mechanism can account for the full repertoire of human memory strategies, as the optimal approach may be task-dependent. For example, human memory experts widely apply the ''memory palace'' strategy, which is empirically better but not captured by temporal context models. Here we show that neural networks optimized for free recall develop diverse retrieval strategies, with only some of them resembling temporal context models.The best-performing models discovered a stimulus-invariant index code that emphasizes the studied position of each list item, instead of its temporal context. This creates a stable scaffold for forward recall akin to the memory palace technique. This index code was more likely to emerge when networks were i) encouraged to recall all studied items rather than prioritizing a few items, and ii) prevented from relying on recency, resonating with human data. Our findings demonstrate that human-like recall patterns can arise from multiple distinct computational mechanisms, and that sequential retrieval using item index is an optimal strategy that explains expert-level recall performance.
Perez, O. D.; Cancino, N.; Hermosilla, D.; Soto, F. A.; Vogel, E. H.
Show abstract
In animal learning research, learning is often represented by plotting a behavioral measure as a function of training trials. A particularly clear case is habituation, a basic form of learning in which repeated presentation of a stimulus produces a decrement in responding. Although retention tests provide the strongest basis for evaluating durable habituation once short-lived performance effects have dissipated, the pattern of response change across stimulus repetitions, or habituation curve, remains theoretically and empirically relevant because it is used to characterize determinants of habituation, individual and clinical profiles, and functional forms, including linear, curvilinear, asymptotic, and mixed incremental-decremental patterns of responding. However, group averaged curves may conceal substantial individual heterogeneity. Here, we analyzed archived human eyeblink habituation data from 157 participants to ask whether the curve shape selected for the group average reflects the curve shapes observed at the individual level. Five candidate functions were fitted separately to each participant and to the corresponding group average. No single function characterized most individuals. More importantly, the model selected for the group average differed from the most frequent individual model in all four groups. When data were pooled across groups, the average favored a dual-process form, a shape that matched the individual plurality in none of them. Simulation analyses showed that averaging heterogeneous individual trajectories can itself produce a group curve that favors a more complex model. Our findings show that group averaged habituation curves should not be treated as direct descriptions of the typical individual trajectory.
Chen, S.; Mueller, H. J.; Shi, Z.
Show abstract
Attentional control balances proactive suppression of predictable distractors with reactive suppression of unexpected ones. Yet, how internal states such as alertness shape this balance is unclear. Using pupillometry and eye tracking across two probability-cueing experiments (conducted in 2024) with varying distractor prevalence, we distinguished tonic (baseline pupil size across blocks) from trial-level pupil size fluctuations (trial-by-trial residual variability in pre-stimulus pupil size). With moderate prevalence, suppression of frequent-region distractors developed gradually, whereas high prevalence induced near-immediate suppression. Behavioral measures (e.g., reaction times) were closely linked to tonic and trial-level pupil size fluctuations. Critically, both alertness components jointly influenced control: during early learning, heightened trial-level pupil size increased distractor capture and reduced target fixations, whereas later on, suppression shifted to a proactive mode resilient to trial-level fluctuations. Under high prevalence, this shift occurred faster. Notably, higher trial-level pupil size generally accelerated first target selection. These findings show that tonic alertness and trial-level alertness fluctuations dynamically regulate reactive and proactive control during statistical learning. Impact StatementThis study shows that people become better at ignoring predictable distractions over time, but that this improvement depends not only on what they have learned about the task environment, but also on their current level of alertness. By combining eye tracking and pupil measures, we found that temporary increases in alertness can sometimes help people orient more quickly to relevant information, yet during earlier stages of learning they can also make attention more vulnerable to distracting events. These findings suggest that successful focus in complex environments depends on a dynamic interplay between learned expectations and moment-to-moment fluctuations in mental state, with implications for understanding sustained attention in settings such as monitoring, driving, and other tasks that require people to stay engaged while resisting distraction.
Zaid, H.; Schaffer, E. S.
Show abstract
In many brain regions, the stimulus tuning of neurons is stable on a timescale of hours but not on a timescale of weeks, a phenomenon often called representational drift. This would seem to imply that these brain regions cannot be used for stable recognition of sensory stimuli or the retrieval of associative memories learned several weeks prior. However, decoding approaches have demonstrated that in some cases, stable decoding of drifting representations is possible. In principle, adaptive decoding provides a plausible resolution to the paradox of how the brain operates with drifting representations, but we lack a deep understanding of what the requirements are for stable decoding to be possible. Here, we offer a general mathematical framework that explains when and why stable decoding from a drifting representation can be achieved. First, we demonstrate that both feedforward and recurrent networks preserve the geometry of their inputs when the network is sufficiently large, meaning that representational drift must also preserve geometry in these networks. Second, we demonstrate that drifting representations that have stable geometry are decodable with adaptive decoders. Therefore, not only the existence of preserved geometry in the presence of representational drift but also the ability to decode from drifting representations simply requires the population of neurons exhibiting representational drift to be large. This theoretical framework not only suggests that preserved geometry should be a general feature of drifting representations, it also explains the conditions under which empirical efforts to measure stable geometry will be successful.
Perez-Bellido, A.; Moreno-Bote, R.; Fuentemilla, L.
Show abstract
Humans exhibit a pervasive drive toward self-consistency, often failing to revise previous decisions even when confronted with contradictory evidence. Here, we investigate the computational mechanisms underlying decision revision in perceptual tasks, examining the regulatory role of metacognition. To do so, we capitalize on a novel paradigm in which participants are repeatedly presented with identical sensory information and allowed to revise their choices after each exposure. Our results reveal that repeated exposure to the same stimulus systematically biases subsequent judgments toward prior responses. Using drift-diffusion modeling, we tested competing explanations incorporating different assumptions about how prior choices affect evidence accumulation. Our findings indicate that consistency biases emerge from asymmetric sensory weighting, selectively amplifying information consistent with previous choices--a phenomenon akin to confirmation bias. Crucially, individuals with higher metacognitive skills exhibited weaker confirmatory biases and more flexible integration of repeated sensory information, enabling greater adaptability in decision-making. These findings highlight the continuous nature of perceptual inference and underscore metacognitions pivotal role in mitigating bias and optimizing decision flexibility.
Pesthy, O.; Toth-Faber, E.; Nagy, C.; Nemeth, M.; Janacsek, K.; Nemeth, D.
Show abstract
Children often outperform adults in probabilistic statistical learning tasks, yet the mechanisms underlying this developmental advantage remain poorly understood. Here, we used eye-tracking measures of belief updating to examine how children and adults acquire and update predictions in a probabilistic sequence-learning task. Using the standard (oculomotor) reaction time measure, children showed stronger statistical learning than adults, replicating previous behavioral findings while revealing a more detailed profile of developmental differences in statistical learning. Critically, children updated their predictions more frequently: they were less likely to repeat previous predictions and more likely to shift their expectations in response to new input. Adults, in contrast, showed greater persistence, tending to maintain prior predictions even when those predictions were inconsistent with the underlying statistical structure. Despite these pronounced differences in updating behavior, the processing and use of prediction errors were remarkably similar across age groups. These findings indicate that developmental differences in statistical learning do not primarily arise from how prediction errors are computed, but rather from how prior beliefs and incoming information are weighted during belief updating. Children's enhanced learning may therefore reflect reduced reliance on stable priors and greater sensitivity to current sensory evidence, supporting a more exploratory learning strategy. Adults, by contrast, appear to favor an exploitative strategy that stabilizes existing predictions but reduces flexibility in probabilistic environments. More broadly, the results suggest that developmental changes in statistical learning may reflect age-related differences in how readily learners revise their predictions in response to incoming evidence. By integrating sensitive oculomotor measures with analyses that probe the mechanisms underlying belief updating, the present study provides a more fine-grained account of how predictive learning changes across development and offers a framework for reconciling previously inconsistent developmental findings in statistical learning.
Casco-Rodriguez, J.; Hong, F.; Brainard, D. H.; Feather, J.; Lipshutz, D.
Show abstract
Representations of the same physical stimulus vary between individuals. Characterizing individual differences has practical implications, but is challenging because these representations are not directly observable. Given a model of how representations vary within a population, we propose a Bayesian adaptive procedure for estimating an individual observer's representation from a series of targeted perceptual discrimination judgments. A key component of our approach is using Fisher information to identify stimulus distortions that efficiently differentiate observers in the population. As a proof of concept, we focus on individual differences in color perception and simulate observers with cone fundamentals drawn from an individual colorimetric observer model. We demonstrate that our approach can recover key aspects of a sampled observer's cone fundamentals using simulated three-alternative forced-choice oddity judgments with approximately 500 trials, corresponding to an experimental duration of approximately one hour. Our Bayesian adaptive framework provides a promising and generalizable approach to efficiently link behavioral measurements to individual differences in sensory representations.
Paro, A. N.; Sheikh, B. I.; Stanford, T. R.; Salinas, E.
Show abstract
The ability to orient or attend to sensory events is generally greater in response to visual and auditory cues occurring together than in response to single-modality cues occurring alone. In such cases the perceptual fusion of cross-modal stimuli (multisensory integration) depends on low-level features (e.g., location, intensity) and follows well established principles. However, less is known about multisensory integration mechanisms when behavioral responses are less direct and require top-down control. Here we investigate this in human participants using an urgent multisensory choice task that effectively dissociates stimulus-driven and goal-driven contributions to performance based on their distinct temporal signatures. Task conditions varied the modality of the cues (auditory, visual, or both), their location (left or right), and the rule defining the correct choice (look toward or away from a given cue). When spatially coincident cues were associated with the same response rule ("look away"), we observed multisensory enhancement and performance remained close to a statistical expectation as the choice process unfolded. However, when spatially disparate cues were associated with different rules but the same target, one cue dominated performance and the other produced crossmodal capture, i.e., low-level competition. The results indicate that the efficacy of multisensory integration is dictated by the stimulus-and goal-driven signals produced by each cue, with all four factors rapidly interacting in accordance to the dynamics of spatial attention. Significance StatementAuditory and visual stimuli located near each other in space and time are typically bound into a single sensory percept that draws attention most effectively. However, it is unclear whether such "multisensory integration" occurs during behaviors that go beyond directly attending or orienting to cue stimuli and require top-down control. We investigated this using a novel task design with which stimulus-driven and goal-driven contributions to performance can be accurately identified. We found that multisensory enhancement depends not so much on the complexity of the requested cue-response associations, but rather on the timing and alignment of the stimulus-and goal-driven signals derived from each cue (auditory and visual) -- similar to the way that such signals dictate the allocation of spatial attention.
Sturrock, M.; Shahrezaei, V.
Show abstract
Approximate Bayesian computation sequential Monte Carlo (ABC-SMC) propagates its particles with a perturbation kernel, and with the standard Normal kernel it degrades sharply as the parameter dimension grows, a failure usually attributed to dimension itself. We show instead that it is governed by the quality of the summary statistics, with dimension entering only through a separate and milder mechanism, and that the two must act together for the Normal kernel to break. The first ingredient is covariance overinflation: the kernel covariance, estimated from the particle cloud, overshoots the true posterior covariance by a factor set by information loss in the summary statistics. We derive this overscaling factor in closed form for a Gaussian model with sufficient statistics and show that it stays modest at any dimension, shrinking toward its baseline value as the tolerance tightens; the extreme values seen in practice (of order 103) are a signature of insufficient summaries, not of dimension. The second ingredient is perturbation overconcentration: the normalised Normal step size concentrates around one as the dimension grows, so every proposal overshoots by the same factor. Either ingredient alone is harmless; only their combination breaks the Normal kernel. A Cauchy kernel (multivariate t with one degree of freedom) removes the concentration, keeping a positive acceptance rate under arbitrary overscaling at a bounded worst-case cost of 1.87x in expected squared jump distance. In a Metropolis-Hastings framework we derive closed-form acceptance rates for both kernels that illustrate the advantage of the Cauchy kernel in this limit. A series of full ABC-SMC computational experiments on five problems at d = 12, including a hierarchical gene-expression model, show the Cauchy reducing the sliced Wasserstein distance to the reference posterior by factors of up to 50 with the same simulation budget. Since the summary statistics are commonly insufficient for the models that require ABC, overinflation is structural and the Cauchy perturbation kernel is the right default for problems in higher dimensions.
Gopnarayan, M. N.; Bavard, S.; Stuchly, E.; Gluth, S.
Show abstract
Social decision-making depends on inferring others hidden preferences from observable behavior. Yet it remains unclear how humans combine choices with process cues such as response times and gaze when learning about others in real-time interaction. Here we combine a novel multi-attribute bargaining task with eye-tracking and show that multiple decision-process cues support preference inference. Across 75 buyer-seller dyads, buyers acceptance rates tracked offer utility, rejection speed reflected decision confidence, and first fixations preferentially targeted the highest-weighted attribute. Sellers adapted subsequent offers using choices, response times, and, when available, gaze cues. A hierarchical inference and choice model suggested that sellers balanced expected utility with expected information gain and updated their beliefs in a Bayesian manner. Although gaze access did not improve overall performance, it changed how sellers used attentional information. These findings shed light on how humans infer others hidden preferences from decision dynamics in real-time social interaction.
Andrade, K. D.; Melton, D. L.; Ries, S. K.
Show abstract
Language production requires the coordination of multiple cognitive processes. The ability to anticipate and override a habitual response in favor of a contextually-appropriate response are key subprocesses of cognitive control which enable speakers to communicate effectively. Word retrieval involves the co-activation of semantically related alternatives from which the speaker must select the appropriate target representation. Although cognitive control mechanisms have been proposed to contribute to resolving semantic interference during language production, the nature of these control processes remain unclear. Studies investigating the temporal dynamics of cognitive control during decision making tasks have led to a distinction between two operating processes: proactive control, initiated prior to the occurrence of conflict, and reactive control recruited after conflict is detected. We investigated the roles of proactive and reactive control in resolving interference between competing linguistic representations during word retrieval. We analyzed congruency sequence effects combined with delta-plot distributional analyses to dissociate potential adjustments in proactive versus reactive cognitive control in a picture-naming task manipulating semantic context compared to a minimally-linguistic Stroop-like paradigm. Reaction time distributional properties following semantically related trials revealed the engagement of proactive control in semantic interference resolution during word retrieval in the PWI task. In contrast, reactive inhibitory control was engaged in resolving semantic interference following low conflict trials. This distinction was not present in the minimally-linguistic task, which did not appear to engage adaptive control to the same extent. These findings demonstrate that both proactive and reactive cognitive control mechanisms contribute to language production, and are engaged dynamically, adjusting trial-by-trial to resolve semantic interference during word retrieval. In addition, our study provides important insight into the comparison of language with other cognitive domains and positions linguistic paradigms as being instrumental in the study of cognitive control dynamics.
Dirks, C. E.; Guest, D. R.; Oxenham, A.
Show abstract
Context effects are ubiquitous across sensory systems and reflect a general encoding principle for both simple and complex stimuli. One simple context effect, contraction bias, manifests in two-interval perception tasks as a bias of the perceived magnitude of the first stimulus toward the center of the overall magnitude range. The underlying cause of contraction bias is unclear. One explanation is that a listeners magnitude estimate of the first stimulus is combined with a perceptual anchor, usually the mean stimulus magnitude, biasing it toward the anchor (sensory model). An alternative explanation is that a listeners response criterion shifts, based on the magnitude of the stimulus pair, relative to the mean magnitude of the stimuli range (decision model). Two pitch-discrimination experiments were performed to test these hypotheses in the auditory domain. The first was a forced-choice discrimination task, where listeners were asked to identify the higher or lower tone in a pair. The second was a same-different task where listeners indicated whether or not the two tones in a pair differed in frequency. Contraction bias was observed in the higher-lower discrimination task, even after extensive perceptual training with feedback. In contrast, no contraction bias was observed in the same-different task. Computational models of the sensory and decision hypotheses were fit to data from both experiments. The sensory model captured the pattern of results the higher-lower experiment but erroneously predicted a contraction bias in the same-different task. The decision model produced similar predictions to the sensory model in the higher-lower task but correctly predicted no contraction bias in the same-different task, and produced lower prediction errors and more stable parameter estimates in both paradigms. Overall, the results suggest that the underlying nature of the contraction bias may reflect decision, rather than sensory, biases based on the context.
Chow, J.; Yang, Y.; Laschowski, B.
Show abstract
Inverse reinforcement learning can recover reward functions from observed behavior, but interpreting those rewards remains a fundamental challenge for understanding intelligent behavior and decision-making. To address this challenge, we introduce a novel framework for reward interpretation that combines reward-function analysis, latent mode assignments, and short-history behavioral analysis to infer latent motivations and behavioral dynamics. As a proof-of-concept, we instantiated the framework using switching inverse reinforcement learning on a large-scale dataset of multi-agent social interactions. Our framework interpreted the learned latent modes as cautious and volatile motivational profiles, demonstrating that recovered reward functions can reveal distinct patterns of behavioral dynamics. More broadly, these findings suggest that the proposed framework provides a promising approach for reverse-engineering and interpreting latent rewards underlying intelligent behavior and decision-making.
Kim, H. E.; Darley, J. O.; Landy, M. S.; Chua, R.; Fox, D. J.
Show abstract
The human sensorimotor system is remarkably effective at automatically parsing total movement error into its constituent parts, the error component due to a perturbation, or externally-generated error (EGE), versus the error component due to motor noise, or internally-generated error (IGE). Participants robustly, and implicitly, adapt to minuscule (2{degrees}) EGEs in the form of randomized visuomotor rotations while ignoring identically-sized errors caused by IGE. This error parsing, and its associated perceptual processes, directly contrasts previous work showing that humans must observe rotations that are > 1.5x the standard deviations of their motor variability, or [≥] 4{degrees}, before explicitly reporting their presence. While the combined results suggest a dissociation between perception for action--which allows for precise and automatic error parsing--and perception for conscious detection, this must be inferred across studies using different methodologies. Here, we combined a within-subjects study design and computational modeling to shed light on the principles underlying implicit adaptation to a perturbation and explicit perturbation detection. Neuro-typical adults participated in two experiments consisting of pseudo-randomized rotations during reaches to a single target, with one session requiring explicit reports after each reach of whether a perturbation was detected. Participants demonstrated a clear dissociation between implicit responses to a perturbation and explicit detection, with robust adaptation to 1{degrees} EGEs, but an inability to reliably report the presence of an EGE until it reached [~] 4{degrees}. For the adaptation task, a model that assumes the participant compares proprioceptive and visual cues to detect a perturbation and corrects for a proportion of this error best fit the data. For signal-detection, a Bayesian causal-inference model in which sensory cues are optimally integrated with a prior on their cause best fit those data. These results indicate that implicit adaptation is dissociated from explicit perturbation detection and the sensorimotor system applies distinct computational strategies to these behaviors.
Riveland, R.; Pouget, A.; Latham, P.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWThere is a gap between neuroscientific theories of learning and the speed of learning observed in many experiments. Since the Cognitive Revolution of the 1950s, compositionality has played a central role in efforts to bridge this gap. Roughly, a compositional system is one where distinct modules are combined according to a set of rules in order to accomplish complex tasks. Recently, significant progress has been made in understanding the emergence of modules in both biological and artificial neural systems. How, and under what conditions, the rules of module recombination are represented in these systems remains an open question. Here we present a neural model that can leverage these rules to dramatically speed up learning. We first show that when faced with multiple tasks which share subcomponents, models learn a low-dimensional representation that captures how subcomponents are reused across the task set. These low-dimensional spaces encode the structure that governs how modules should be recombined. Restricting learning to these subspaces greatly reduces the amount of experience needed to acquire a novel task, even when learning from reinforcement on single trials. In some cases, we can leverage the geometric regularities of these representations to reduce learning to a form of hypothesis testing over a small set of discrete points. Finally, we use this theory to model both behavioral and neural data from non-human primates performing a compositional task, and show that key features in this data are consistent with a model in which exploration during learning is restricted to these low-dimensional spaces. Overall, this work shows that the advantages of modularity in neural systems can be greatly improved upon when models represent the structure of module reuse. Both these features working in tandem lead to learning on timescales similar to biological intelligences, and hence provide a model for how such fast, adaptable behavior can emerge from systems of neurons.
XU, M.; REN, Y.
Show abstract
Building upon foundational psychological theories of event segmentation, this study addresses the limitation of overreliance on temporal boundaries as the primary segmentation criterion. Drawing on two experiments of direct and indirect causation in Mandarin Chinese, this study demonstrates how cognitive segmentation granularity and semantic integration jointly shape syntactic encoding. Results reveal distinct event encoding patterns for direct and indirect causation: coarse-grained segmentation leads to compact syntactic structures (e.g., verb-resultatives), while fine-grained segmentation yields varied multi-clausal expressions. Chinese speakers update event models via prediction errors of intentionality and protagonists, and tend to establish event boundaries at goal-relevant action endpoints when construing causal chains. These conceptual dimensions exert a modulating influence on both event segmentation and semantic integration. We propose a triad model integrating event segmentation, semantic integration, and linguistic specificity, providing a unified framework for elucidating the mind-language interface in conceptual construction and event coding of causation.
Lloyd, B.; Kikumoto, A.; Wurm, F.; Vives, M.-L.
Show abstract
Learning is typically understood as a process driven by prediction errors, when outcomes differ from expectations. Yet it remains unclear whether outcomes that perfectly match expectations are psychologically and computationally meaningful. Here, we tested whether zero prediction errors shape affect, belief updating, and neural feedback processing in human reinforcement learning. Participants repeatedly predicted rewards in environments varying in uncertainty, with a subset of trial outcomes manipulated to exactly match their predictions. Zero prediction errors produced the highest momentary happiness, and computational modeling showed that behavior was best explained by a model in which zero prediction errors induce a distinct latent belief state that guides subsequent updating, particularly under higher uncertainty and in individuals with greater intolerance of uncertainty. Outcome-locked EEG analyses further showed that zero prediction errors elicited distinct P3-like responses, with residual neural activity predicting attenuated updating after zero prediction errors but enhanced updating after standard prediction errors. These findings suggest that perfect predictions are not neutral, but informative events that actively shape affect, behavior, and neural feedback processing.
Mahajan, P.; Seymour, B.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWDopamine activity in the tail of the striatum (TS) presents a novel challenge for reinforcement-learning theories of dopamine. Some studies suggest that TS-projecting dopamine signals encode aversive or threat prediction errors, whereas others argue that they encode action prediction errors involved in soft-habit formation. Here, we show that these accounts need not be mutually exclusive. We instantiate an entropy-regularised reinforcement-learning model in which TS-projecting dopamine neurons update both aversive values and the default policy. In this model, threat belief gates aversive value initialisations, producing TS-like activity during retreat from potentially threatening novel objects, while default-policy learning generates action prediction error signals that decline as actions become habitual. Our results further suggest why both of these signals may need to coexist in the temporal difference errors in our model, and qualitatively reproduce key response patterns and simulations from studies previously used to support both views. Beyond this descriptive reconciliation, our model simulations also highlight the normative role of the tail of the striatum in cautious behaviours in the context of potential threats and stable learning in the face of outcome uncertainty.